Papers with human evaluation
Copied to clipboard
| Challenge: | Existing models for visual entailment and visual question-answering have limited ability to understand figurative meaning in images and captions. |
| Approach: | They propose a task framing the figurative meaning understanding problem as an explainable visual entailment task where the model has to predict whether the image entitles a caption and justify the predicted label with a textual explanation. |
| Outcome: | The proposed dataset contains 6,027 image, caption, label, explanation instances covering five diverse figurative phenomena. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have revolutionized natural language processing with impressive performance across various tasks. |
| Approach: | They propose a framework for automated evaluations of large language models . they open-source their code at https://github.com/WisdomShell/FreeEval . |
| Outcome: | The framework is open-source and can be used to develop and validate new evaluation methods. |
Copied to clipboard
| Challenge: | Neural models generate the most common and generic responses all the time . Empirical results show that our method can significantly improve the diversity of responses generated by sequence-to-sequence models. |
| Approach: | They propose an iterative training process and ensemble method based on boosting to improve the diversity of responses generated by neural models. |
| Outcome: | Empirical results show that the proposed method significantly improves diversity and relevance of responses generated by all models. |
Copied to clipboard
| Challenge: | Existing systems or studies lack interactivity and do not provide off-the-shelf signals. |
| Approach: | They propose an interactive system that extracts and highlights crucial financial signals . they integrate pre-trained BERT representations and a fine-tuned BERT highlighting model . |
| Outcome: | The proposed system extracts and highlights key financial signals efficiently and precisely. |
Copied to clipboard
| Challenge: | a recent study shows that sentiment analysis datasets lack context in which an opinion was expressed and are limited by a few emotion categories. |
| Approach: | They propose to ground an LLM-based model into a corpus of narratives to generate stories-character-centered utterances with unique contexts over 28 emotion classes. |
| Outcome: | The proposed model generates non-repetitive story-character-centered utterances with unique contexts over 28 emotion classes. |
Copied to clipboard
| Challenge: | Code-switching (CSW) text generation is a popular solution to address data scarcity. |
| Approach: | They compare linguistic theories, lexical replacements and back-translation approaches to Egyptian Arabic-English CSW. |
| Outcome: | The proposed methods perform best on machine translation and quality evaluation. |
Copied to clipboard
| Challenge: | Using back-translation, we can improve generalization by using noisy channel re-ranking and ensembling. |
| Approach: | They propose to use BPE-based transformer models to leverage monolingual data to improve generalization and use noisy channel re-ranking and ensembling to improve results. |
| Outcome: | The proposed system improves on the baseline system trained exclusively on the provided small parallel dataset, and the human evaluation and BLEU score are higher. |
Copied to clipboard
| Challenge: | Ambiguity is a linguistic tool for encoding information efficiently, yet it also causes misunderstandings and disagreements. |
| Approach: | They propose a constrained generation task for explaining ambiguous claims in fact-checking by editing them to spell out an interpretation that can be unequivocally supported by the given evidence. |
| Outcome: | The proposed model disambiguates claims 72% of the time compared to a simple copy baseline and a Large Language Model baseline. |
Copied to clipboard
| Challenge: | Our study examines ChatGPT’s performance on Arabic languages and dialectal varieties. |
| Approach: | They conduct a large-scale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets. |
| Outcome: | The proposed model outperforms smaller models on Arabic dialects compared to GPT-4's Modern Standard Arabic and Dialectal Arabic (DA) |
Copied to clipboard
| Challenge: | Prior studies modeled multimodal UI grounding in one round, but such an interaction is inherently iterative. |
| Approach: | They propose a task where a user and an agent collaborate on an interface screen . they use a dataset of 77,820 sequences of human user-agent interaction on mobile interfaces . |
| Outcome: | The proposed task improves the absolute task completion by 18% over the entire test set and 31% over the challenging split. |
Copied to clipboard
| Challenge: | Existing systems require developers to manually generate and annotate a large number of utterances. |
| Approach: | They propose a system that guides ordinary software developers to build a high quality NLU engine from scratch. |
| Outcome: | The proposed system shows that iterative pruning of incorrect utterances reduces human workload and cognitive load. |
Copied to clipboard
| Challenge: | Text detoxification can mitigate the harms of toxicity by rephrasing text to remove offensive meaning, but subtle toxicity remains challenging to tackle. |
| Approach: | They propose a text detoxification algorithm that combines controllable generation and text rewriting methods using a Product of Experts and autoencoder language models to find candidate words to mask and potentially replace. |
| Outcome: | The proposed method outperforms baselines on automatic metrics and is preferred 2.1 times more in human evaluation. |
Copied to clipboard
| Challenge: | a clinical note is a document that documents a doctor's interaction with a patient . authors show that LLMs can be used to measure quality indicators . |
| Approach: | They analyze two different approaches to generate different sections of a SOAP note . they use PEGASUS-X Transformer models to examine note consistency . |
| Outcome: | The proposed approach leads to similar ROUGE values and no difference in Factuality metric . human reviewers perform the same tasks with roughly the same agreement as the LLMs . |
Copied to clipboard
| Challenge: | Existing document translation pipelines face a tension between linguistic processing and layout preservation. |
| Approach: | They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content. |
| Outcome: | The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision. |
Copied to clipboard
| Challenge: | Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems . |
| Approach: | They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language. |
| Outcome: | The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature. |
Copied to clipboard
| Challenge: | Existing machine translation metrics have poor correlations with human assessments . entropy-based evaluations are often limited to a limited number of samples . |
| Approach: | They propose a fast and unsupervised approach to enhance machine translation metrics using entropy by introducing sentence-level difficulty. |
| Outcome: | The proposed method outperforms existing metrics on five sub-tracks in the WMT19 Metrics shared tasks. |
Copied to clipboard
| Challenge: | a recent study shows that human evaluations of dialogue systems weakly reflect human judgments. |
| Approach: | They propose a BERT-based model that fine-tunes a model with three prediction heads to predict whether the system-generated output is natural, fluent, and informative. |
| Outcome: | The proposed model achieves an average accuracy of 77% over the 3 labels . it also uses three different models to compute the labels compared to three separate models . |
Copied to clipboard
| Challenge: | Neural models for text generation are often designed in an end-to-end fashion, limiting their practical usability in downstream applications. |
| Approach: | They propose a method to compute image representations specific to each sentential context and exploiting diverse sentence states to ensure topical continuity and content diversity of generated radiology reports. |
| Outcome: | The proposed method outperforms baselines on objective metrics and human evaluations by 18% and 29% respectively in the evaluation for informativeness and content ordering respectively. |
Copied to clipboard
| Challenge: | Existing paraphrase datasets are mainly from news, novels, or social media platforms. |
| Approach: | They propose to build a large-scale paraphrase dataset using intra-paper and inter-paper methods . they use PDBERT as a general paraphrase discovering method to take advantage of paraphrased sentences . |
| Outcome: | The proposed dataset includes 33,981 paraphrase pairs from ACL and 316,063 pairs from arXiv . the major advantages of paraphrases lie in the prominent length and textual diversity . |
Copied to clipboard
| Challenge: | Existing tools for evaluation of translation models focus on high-level metrics like BLEU or COMET scores, which are time-consuming and prone to error. |
| Approach: | They propose a toolkit that provides a detailed analysis of translation models and a user-friendly interface. |
| Outcome: | The toolkit shows superior performance over COMET and SacreBLEU packages under enjoybility and understandbility criteria. |
Copied to clipboard
| Challenge: | Medical images are widely used in clinical decision-making, where writing radiology reports can be enhanced by automatic solutions to alleviate physicians’ workload. |
| Approach: | They propose an approach with reinforcement learning over a cross-modal memory to better align visual and textual features for radiology report generation. |
| Outcome: | The proposed approach improves cross-modal alignment on two English radiology report datasets and human evaluation confirms the results. |
Copied to clipboard
| Challenge: | sarcasm generators assume intended meaning is opposite of literal meaning . sarcastically generated responses are more specific and coherent to input . |
| Approach: | They propose a system that generates sarcastic responses to a given utterance . they ground their generation process on a formal theory that unambiguously differentiates . |
| Outcome: | The proposed system generates sarcastic responses to a given utterance. |
Copied to clipboard
| Challenge: | MEET-MR provides a comprehensive benchmark for evaluating English–Thai machine translation systems. |
| Approach: | They propose a benchmark for evaluating English–Thai machine translation systems . they use the Multidimensional Quality Metrics framework to provide fine-grained human judgements of translation quality. |
| Outcome: | The dataset covers nine domains providing linguistic and contextual diversity. |
Copied to clipboard
| Challenge: | Using large language models, agents can assist with natural language tasks when given access to confidential data. |
| Approach: | They created a synthetic dataset consisting of confidentiality-aware planning and deduction tasks in organizational access control. |
| Outcome: | The proposed model can perform tasks similar to humans when given access to confidential data. |
Copied to clipboard
| Challenge: | Existing studies on style transfer for text are lacking a standard set of evaluation practices. |
| Approach: | They propose a set of metrics for automated evaluation that are more strongly correlated with human judgment and show tradeoffs between aspects of interest. |
| Outcome: | The proposed models exhibit tradeoffs between aspects of interest and human judgment, demonstrating the importance of evaluating them at specific points of their tradeoff plots. |
Copied to clipboard
| Challenge: | Existing methods for text summarization evaluation do not correlate well with human judgments . evaluators that use Likert scale scores are limited in their ability to perform deeper analysis. |
| Approach: | They propose a fine-grained evaluator specifically tailored for the summarization task using large language models. |
| Outcome: | The proposed method improves on open-source and proprietary LLMs and shows better completeness and conciseness than existing methods. |
Copied to clipboard
| Challenge: | End-to-end dialogue systems with monolithic neural architecture are often trained with input-output utterances without taking into account the entire annotations available in the corpus. |
| Approach: | They propose an end-to-end neural architecture for goal-oriented dialogue systems that addresses both challenges . they propose a modular architecture where modules are optimized individually . |
| Outcome: | The proposed system achieved the top position in the human evaluation task . it is based on a neural architecture that can be integrated with external systems . |
Copied to clipboard
| Challenge: | generative pre-trained models face challenges on constrained writing tasks like poem generation . brian mccartney: BIPro improves the zero-shot generation quality on constricted writing tasks . |
| Approach: | They propose a framework that leverages two block inverse prompting methods to improve the quality of constrained writing tasks. |
| Outcome: | BIPro significantly improves the quality of Chinese poem generation without priming or training. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models are useful, honest, harmless (HHH) however, RLHF requires high hardware resources and human efforts. |
| Approach: | They propose a framework that allows LLMs to align themselves with HHH . they use IF and reinforcement learning from human feedback to fine-tune their models . |
| Outcome: | The proposed framework achieves similar performance to RLHF and human-generated models with a minimal alignment tax. |
Copied to clipboard
| Challenge: | This thesis examines how humans and models perceive writing style under controlled perturbations. |
| Approach: | They examine how humans and models perceive writing style under controlled perturbations . they also examine whether perturbations that reduce algorithmic recognition obscure stylistic identity . |
| Outcome: | The proposed research compares models and humans to find out how linguistic cues affect writing style . it will clarify how linguistic cue contributes differently to human and algorithmic perception of style - a cnn.com article argues . |
Copied to clipboard
| Challenge: | In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon the insights from disciplines such as user experience research and human behavioral psychology to ensure that the results are reliable. |
| Approach: | They propose a framework for human evaluation of generative large language models that takes into account usability, aesthetics and cognitive biases. |
| Outcome: | The proposed framework is based on the framework proposed by Deutsch and alnajjar . it is aimed at ensuring that human evaluation is accurate in the age of generative AI . |
Copied to clipboard
| Challenge: | Dialogue systems for interaction with humans are becoming more popular . the best way to estimate their success is through means of human evaluation . |
| Approach: | They investigate the effectiveness of perceiving dialogue evaluation as an anomaly detection task. |
| Outcome: | The proposed approach is based on four models and shows negative results . the proposed approach could be used in the future to improve human-led dialogue evaluations. |
Copied to clipboard
| Challenge: | In general, speech synthesis for Indigenous languages is underdeveloped compared to the majority of languages. |
| Approach: | They propose to train a multilingual model on three typologically similar languages to improve performance over monolingual models. |
| Outcome: | The proposed model can train on three similar languages with high performance and is highly competitive with self-attention architectures with higher memory efficiency. |
Copied to clipboard
| Challenge: | a new method to extract user attributes from dialogues is needed to improve user understanding. |
| Approach: | They propose to leverage dialogues with conversational agents to automatically extract user attributes from dialogues. |
| Outcome: | The proposed model surpasses retrieval and generation baselines on human evaluation. |
Copied to clipboard
| Challenge: | Existing models that ground knowledge and persona at the same time are limited, leading to hallucination and a passive way of using personas. |
| Approach: | They propose a conversational agent that grounds external knowledge and persona simultaneously and a retrieval augmented generation model that generates utterances with lesser hallucination and more engagingness. |
| Outcome: | The proposed agent generates the utterance with lesser hallucination and more engagingness utilizing retrieval augmented generation with knowledge-persona enhanced query. |
Copied to clipboard
| Challenge: | Neural text generation (data- or text-to-text) demonstrates remarkable performance when training data is abundant which for many applications is not the case. |
| Approach: | They propose a technique to treat hallucinations as a controllable aspect of the generated text without dismissing any input and without modifying the model architecture. |
| Outcome: | The proposed technique can be used on a WikiBio dataset and in a human evaluation. |
Copied to clipboard
| Challenge: | Question Under Discussion (QUD) uses implicit questions to reveal discourse relationships between sentences. |
| Approach: | They propose a framework that selectively decodes the QUD dependency structures considering the QUC criteria. |
| Outcome: | The proposed framework outperforms the state-of-the-art baseline models by 9% in human evaluation and 4% in automatic evaluation. |
Copied to clipboard
| Challenge: | Multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages. |
| Approach: | They propose a method where a reward model is trained on preference data in one source language and applied to other target languages. |
| Outcome: | The proposed approach is effective under comprehensive evaluation settings, including human evaluation. |
Copied to clipboard
| Challenge: | et al., 2018) show that human raters prefer corrected translations over the baseline ones. |
| Approach: | They propose a monolingual model to correct inconsistencies between sentences . they use monolingual document-level data to train the model . |
| Outcome: | The proposed model improves translations of contextual phenomena in English-Russian translation task. |
Copied to clipboard
| Challenge: | Mixed initiative dialogue systems allow all interacting agents to initiate actions to control the interaction. |
| Approach: | They propose to prompt large language models as a drop-in replacement for fine-tuning on conditional generation. |
| Outcome: | The proposed prompts improve fine-tuning and ground truth responses . the results show that generated responses are high . |
Copied to clipboard
| Challenge: | Existing methods for controlling coarse attributes are less effective for finer-grained attributes and suffer from inefficiencies when many attributes must be handled jointly. |
| Approach: | They propose a controlled text generation model that allows fine-grained control over a large number of real-valued linguistic attributes. |
| Outcome: | The proposed model achieves the lowest average control error among evaluated methods while remaining efficient at inference and receiving the highest fluency scores in human evaluation. |
Copied to clipboard
| Challenge: | a large study of machine translation systems shows poor evaluation procedures can lead to erroneous conclusions. |
| Approach: | They propose an evaluation methodology grounded in explicit error analysis based on the Multidimensional Quality Metrics framework. |
| Outcome: | The proposed evaluation methodology outperforms crowd workers in two languages . it shows that human-based metrics outperformed crowd workers . |
Copied to clipboard
| Challenge: | The Annals of Joseon Dynasty contain the daily records of the Kings of Joseont, the 500-year kingdom preceding the modern nation of Korea. |
| Approach: | They propose a neural machine translation model that translates historical documents written in Hanja to more easily understandable Korean and to English. |
| Outcome: | The proposed model outperforms baseline models in terms of BLEU scores for both contemporary Korean and English translations. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics are based on procedures that diverge from human evaluation. |
| Approach: | They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system. |
| Outcome: | The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics. |
Copied to clipboard
| Challenge: | Literature in Natural Language Processing (NLP) typically labels whole language with strict type of morphology, e.g. fusional or agglutinative. |
| Approach: | They propose to quantify morphological typology at the word and segment level by using two indices: synthesis (e.g. analytic to polysynthetic) and fusion (agglutinative to fusional). |
| Outcome: | The proposed method reduces the rigidity of NLP classification claims by measuring morphological diversity at the word and segment level. |
Copied to clipboard
| Challenge: | a key ingredient of neural machine translation is the use of large datasets with different but consistent translation styles . however, the models do not capture the variety of translators' styles from the data . a recent study shows that style-augmented models can capture the style variations of translator . |
| Approach: | They propose to augment a neural machine translation model with translator information . they use TED talk datasets to model and control translator-related stylistic variations . |
| Outcome: | The proposed models capture the style variations of translators and generate translations with different styles on new data. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have great potential for synthetic data generation. |
| Approach: | They show that large language models can generate useful data even for complex tasks . they use a symmetric task difficulty asymmetry to prompt an LLM to generate plausible input text for a target output structure. |
| Outcome: | The proposed approach outperforms existing models by a substantial margin on closed information extraction tasks with 1.8M data points and 770M parameters. |
Copied to clipboard
| Challenge: | Text-to-SQL benchmarks are used to evaluate progress made in the field . however, matching a model-generated SQL query to a reference SQL query fails due to various reasons. |
| Approach: | They conduct an extensive evaluation of text-to-SQL benchmarks and re-evaluate some of the top-performing models. |
| Outcome: | The results show that a recent model surpasses the gold standard reference queries in the Spider benchmark in human evaluation. |
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
Copied to clipboard
| Challenge: | CaseSumm is a dataset for long-context summarization in the legal domain . human groundtruth summaries are often not available for legal summarizing . |
| Approach: | They propose a dataset for long-context summarization that includes SCOTUS opinions and their official summaries. |
| Outcome: | The proposed dataset is the largest open legal case summarization dataset . it outperforms larger models on automatic metrics and human evaluation . |
Copied to clipboard
| Challenge: | Existing benchmarks for audio-centric interaction have impeded advancements in this field . AIR-Bench evaluates LALMs' ability to understand audio signals and interact with humans . |
| Approach: | They propose a benchmark to evaluate the ability of large audio-language models to understand audio signals . they use 19 tasks with approximately 19k single-choice questions to examine single-task ability . |
| Outcome: | The proposed framework evaluates the ability of large audio-language models to understand audio signals and interact with humans in the textual format. |
Copied to clipboard
| Challenge: | under the pandemic of COVID-19, people experiencing COVI D19-related symptoms have a pressing need to consult doctors. |
| Approach: | They develop a medical dialog system that can provide COVID19-related consultations . they use two dialog datasets containing conversations between doctors and patients . |
| Outcome: | The proposed system can provide COVID19-related consultations, but is too small compared with general-domain dialog datasets. |
Copied to clipboard
| Challenge: | Existing methods for generating text are unsupervised and require supervision. |
| Approach: | They propose an unsupervised method that uses two off-the-shelf pretrained LMs in opposite directions to apply them to non-sequential tasks. |
| Outcome: | The proposed method outperforms strong unsupervised baselines on paraphrasing and abductive text infilling. |
Copied to clipboard
| Challenge: | Existing classification-based models are poorly per-form for tail labels and ignore semantic relations among labels. |
| Approach: | They propose to guide label generation using label cluster information to hierarchically generate lower-level labels. |
| Outcome: | The proposed model outperforms classification and generation baselines on tail labels and improves in four popular XMC benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for code retrieval struggle to balance scalability and annotation quality. |
| Approach: | They propose a method that integrates functions called within the repository and information on third-party APIs to enhance the annotation context. |
| Outcome: | The proposed method improves the annotation context by incorporating functions called within the repository and information on third-party API functionalities. |
Copied to clipboard
| Challenge: | Existing dialogue systems focus on functional goals, open-domain chatbots on socially engaging conversations. |
| Approach: | They propose to add chit-chat to ENhance Task-ORiented dialogues by a human-assisted data collection approach to augment task-oriented dialogues with minimal annotation effort. |
| Outcome: | The proposed models can code-switch between task and chit-chat to be more engaging, interesting, knowledgeable, and humanlike while maintaining competitive task performance. |
Copied to clipboard
| Challenge: | et al., 2018a): a poor phrasing may make the conversation go awry. |
| Approach: | They propose a model that can help suggest rephrasings of toxic comments in a more civil manner. |
| Outcome: | The proposed model generates sentences that are more fluent and better at preserving the initial content compared to earlier systems and human evaluation. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for natural language generation are dominated by similarity-based metrics. |
| Approach: | They propose a multi-dimensional evaluator for natural language generation that integrates multiple dimensions into one evaluer. |
| Outcome: | The proposed evaluator improves on three typical NLG tasks and improves with external knowledge. |
Copied to clipboard
| Challenge: | FreeTalky is a deep learning-based foreign language learning platform for people who experience anxiety dealing with foreign languages. |
| Approach: | They propose a deep learning-based foreign language learning platform called FreeTalky . it employs a humanoid robot NAO and various deep learning models . |
| Outcome: | The proposed system provides personalized learning based on persona dialogue and grammar error correction, and also helps alleviate xenoglossophobia by replacing the real human in the conversation with a NAO robot, through human evaluation. |
Copied to clipboard
| Challenge: | Existing evaluation methods for text style transfer are unsatisfactory. |
| Approach: | They propose to use a graph-based method to extract attribute content from sentences . they propose an efficient regularization to leverage attribute-dependent content as guiding signals. |
| Outcome: | The proposed method is based on a YELP and IMDB dataset and it is able to detect errors in the human evaluation. |
Copied to clipboard
| Challenge: | Existing methods for evaluating progress in natural language generation tasks are expensive, difficult to reproduce, and non-reusable. |
| Approach: | They propose a new automatic evaluation method for NLG called Near-Negative Distinction that repurposes prior human annotations into NND tests. |
| Outcome: | The proposed method achieves higher correlation with human judgments than standard NLG evaluation metrics. |
Copied to clipboard
| Challenge: | Currently, large language models (LLMs) based on Open domain Natural language planning have limited application potential. |
| Approach: | They propose a dataset with a baseline for Open domain Natural language planning . the dataset provides the largest dataset for textual procedures to date . |
| Outcome: | The proposed dataset provides the largest dataset for textual procedures to date . it leverages entity-attribute-level action models to reveal relevant physical properties . |
Copied to clipboard
| Challenge: | Existing approaches to formalizing mathematical statements face limitations in accuracy, especially in the context of complex, highlevel problems that involve sophisticated mathematical reasoning. |
| Approach: | They propose a CriticLean framework that elevates the role of the critic from a passive validator to an active learning component and introduce a benchmark to measure models’ ability to distinguish semantically correct from incorrect formalizations. |
| Outcome: | The proposed framework outperforms open- and closed-source benchmarks and shows that it significantly outperformed existing models. |
Copied to clipboard
| Challenge: | a novel framework for conversational question answering from unlabeled documents has been proposed . a large-scale dataset of synthetic conversations is available for use in real-world applications . |
| Approach: | They propose a framework for conversational question answering from unlabeled documents . they propose 'SimSeek' framework that simulates conversation from unlabelled documents based on two scenarios . |
| Outcome: | The proposed framework achieves state-of-the-art performance on a recent CQA benchmark, QuAC. |
Copied to clipboard
| Challenge: | Existing sentence compression methods do not handle syntactic features, causing performance degradation . et al. (2015) reported that the longer the input sentences are, the worse the performance becomes. |
| Approach: | They propose a higher-order syntactic attention network that handles higher-level dependency features as an attention distribution on LSTM hidden states. |
| Outcome: | The proposed method outperforms baseline methods on a Google sentence compression dataset. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have limitations in grounding ideas and mitigating confirmation bias during refinement. |
| Approach: | They propose a framework that integrates a Motivational Knowledge Graph with a Q-Driven Socratic Ideator to enhance LLM ideation. |
| Outcome: | The proposed framework enhances LLM ideation by integrating a Motivational Knowledge Graph with a Q-Driven Socratic Ideator. |
Copied to clipboard
| Challenge: | Previous work using adversarial methods has struggled to produce high-quality outputs. |
| Approach: | They propose a method that transforms a sentence to alter a specific attribute while preserving its attribute-independent content. |
| Outcome: | The proposed method generates grammatical and appropriate responses on 22% more inputs than the best previous system, averaged over three attribute transfer datasets. |
Copied to clipboard
| Challenge: | In this paper, we explore creative generation with a focus on puns. |
| Approach: | They propose an unsupervised approach to generating puns using lots of raw text and a surprisal principle. |
| Outcome: | The proposed approach generates puns 30% of the time, doubles the neural generation baseline. |
Copied to clipboard
| Challenge: | Existing models of paraphrase generation are based on a syntactic sketch, but prior work has included inductive bias. |
| Approach: | They propose a method for learning decompositions of dense encodings as a sequence of discrete latent variables that make iterative refinements of increasing granularity. |
| Outcome: | The proposed model improves on human paraphrase generation by predicting syntactic sketches at test time. |
Copied to clipboard
| Challenge: | Existing methods to generate source code summaries are coarse-grained and noise-filled . however, they do not capture contextual code semantics and are often outdated in continuous software iteration. |
| Approach: | They propose a fine-grained Token-level retrieval-augmented mechanism on the decoder side to enhance performance of neural models. |
| Outcome: | The proposed method produces more low-frequency tokens and is interpretable. |
Copied to clipboard
| Challenge: | Existing methods for fine-grained text sentiment transfer only reverse the sentiment polarity of text, but they lack a robust and parallel learning algorithm. |
| Approach: | They propose a novel fine-grained text sentiment transfer task that revises a sequence to satisfy a given sentiment intensity while preserving the original semantic content. |
| Outcome: | The proposed model outperforms existing methods by a large margin in automatic evaluation and human evaluation. |
Copied to clipboard
| Challenge: | Using neural machine translation to approximate human parity is difficult due to the lack of parallel training corpora. |
| Approach: | They propose an end-to-end deep learning framework for quality estimation and automatic post-editing of machine translation output. |
| Outcome: | The proposed framework achieves state-of-the-art performance on the English–German dataset and human translators can significantly expedite their post-editing processing with the model. |
Copied to clipboard
| Challenge: | Experimental results show that our model can achieve a significant improvement in terms of metric-based evaluation and human evaluation compared with the state-of-the-art exposure bias approaches. |
| Approach: | They propose a novel adaptive switching mechanism which automatically transits between ground-truth learning and generated learning regarding the word-level matching score. |
| Outcome: | The proposed model improves on Chinese and English reddit datasets compared with state-of-the-art models on the word-level matching score. |
Copied to clipboard
| Challenge: | Existing evaluation methods for dialogue systems rely on human judges to label quality of generated text. |
| Approach: | They propose a machine learning approach to reduce the effort of human evaluation by learning the human judgment on comparing two generative dialogue systems. |
| Outcome: | The proposed method reduces the effort of human evaluation by learning which generative models is better in each dialog context. |
Copied to clipboard
| Challenge: | Neural machine translation models still face various challenges including fragility and lack of style flexibility. |
| Approach: | They propose to incorporate prompts into neural machine translation to improve translation control and style flexibility. |
| Outcome: | Empirical results show that the proposed method improves translation control and quality and improves human-in-the-loop translation. |
Copied to clipboard
| Challenge: | Existing methods focus on learning a direct mapping from pure code to summaries, overlooking the heterogeneity gap between code and summary. |
| Approach: | They propose a framework that uses chain of comments as auxiliary intermediate information to bridge the gap between code and summaries. |
| Outcome: | The proposed framework outperforms baseline models and multiple code Large Language Models by a large margin. |
Copied to clipboard
| Challenge: | Existing evaluation methods for human-machine interactions are static and can be misleading. |
| Approach: | They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement. |
| Outcome: | The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment. |
Copied to clipboard
| Challenge: | Multi-sentence compression aims to generate a grammatical but reduced compression from multiple input sentences while retaining key information. |
| Approach: | They propose a neural rewriter for multi-sentence compression that does not need any parallel corpus. |
| Outcome: | Empirical studies show that the proposed approach achieves comparable results upon automatic evaluation and improves the grammaticality of compression based on human evaluation. |
Copied to clipboard
| Challenge: | reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable . |
| Approach: | They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable . |
| Outcome: | The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable . |
Copied to clipboard
| Challenge: | Existing studies for summarization evaluation exhibit low inter-annotator agreement or lack scale. |
| Approach: | They propose a modified summarization salience protocol based on fine-grained semantic units and a robust summarizing evaluation benchmark. |
| Outcome: | The proposed protocol is based on fine-grained semantic units and allows for high inter-annotator agreement. |
Copied to clipboard
| Challenge: | Recent studies on open-book QG have achieved promising progress, but generating natural questions under a more practical closed-book setting remains a challenge. |
| Approach: | They propose a QG model that stores more information in its parameters through contrastive learning and an answer reconstruction module. |
| Outcome: | The proposed model outperforms baselines in automatic evaluation and human evaluation on a public dataset and a new WikiCQA dataset. |
Copied to clipboard
| Challenge: | Recent work proposes a method to optimize pipelined dialogue systems by fine-tuning modules directly. |
| Approach: | They propose a new post-processing component for natural language generation (NLG) they use dialogue act contribution to evaluate contribution of GenPPN-generated utterances . |
| Outcome: | The proposed method improves the performance of task-oriented dialogue systems by modifying arbitrary modules including non-differentiable ones. |
Copied to clipboard
| Challenge: | Existing methods to generate adversarial examples for relation classification are vulnerable to adversarials. |
| Approach: | They propose a method that uses most important parts of speech to substitute words with synonyms or hyponyms to generate adversarial texts of high quality. |
| Outcome: | The proposed method can generate adversarial texts of high quality and most relationships can be correctly identified in the process of human evaluation. |
Copied to clipboard
| Challenge: | Using large language models (LLMs) to generate human-like text has raised concerns about misuse, especially in low-resource languages like Urdu. |
| Approach: | They propose a dataset that contains documents, paragraphs, and sentences . they conducted human evaluations and automated evaluations . |
| Outcome: | The proposed dataset shows that distinguishing between human and machine-generated text is challenging for both humans and LLMs. |
Copied to clipboard
| Challenge: | Recent approaches to question generation have used modifications to a Seq2Seq architecture inspired by advances in machine translation. |
| Approach: | They propose to use a Seq2Seq architecture to train models to generate one-step-ahead predictions, but at test time, the model is asked to generate a whole sequence, causing errors to propagate through the generation process. |
| Outcome: | The proposed model is trained to generate a plausible question, conditioned on an input document and answer span within that document. |
Copied to clipboard
| Challenge: | False or misleading narratives spread rapidly on social networks, posing challenges for non-experts in discerning credible information. |
| Approach: | They propose a model for fallacious reasoning that focuses on implicit fallacies between relevant content and the inaccurate claim and requires models to verbalize the fallacious thinking in addition to classifying it. |
| Outcome: | The proposed model focuses on implicit fallacies between relevant content and the inaccurate claim and requires models to verbalize the fallacious reasoning in addition to classifying it. |
Copied to clipboard
| Challenge: | Existing methods to correct factual errors are limited to labeled claims . a recent task of fact verification has attracted significant attention . |
| Approach: | They propose a task of factual error correction that performs edits to a claim so that the generated rewrite is better supported by evidence. |
| Outcome: | The proposed method produces accurate factual error corrections for 5x more instances in human evaluation and a .125 increase in SARI score. |
Copied to clipboard
| Challenge: | Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts . |
| Approach: | They propose to account for belief bias in explainability by using models of varying quality and adversarial examples. |
| Outcome: | The proposed methods show that results change when using models of varying quality and adversarial examples. |
Copied to clipboard
| Challenge: | Large language models (LLMs) provide unprecedented flexibility in defining and executing complex, creative natural language generation tasks. |
| Approach: | They propose a framework that consists of input manipulation, reference data, and output measurement to explore citation text generation. |
| Outcome: | The proposed framework explores citation text generation, a popular scholarly NLP task that lacks consensus on the task definition and evaluation metric and has not yet been tackled within the LLM paradigm. |
Copied to clipboard
| Challenge: | In previous work, a large number of human dialogues are required to train dialogue agents. |
| Approach: | They propose loop-clipping policy optimisation to eliminate useless responses by clipping loops from dialogue history and clipping advantage to distinguish useless actions from others. |
| Outcome: | The proposed method achieves 80% success rate on a Cambridge restaurant dialogue system using 260 training dialogues compared to baseline of 2160 dialogues. |
Copied to clipboard
| Challenge: | Existing methods for knowledge selection focus on relevance between knowledge and dialogue context, ignoring personal preference for knowledge. |
| Approach: | They propose to introduce personal memory into knowledge selection in chatbots to address personalization issue by integrating personal memory and inverse mapping into a closed loop. |
| Outcome: | The proposed method outperforms existing methods significantly on automatic evaluation and human evaluation. |
Copied to clipboard
| Challenge: | Existing metrics for multimodal large language models only focus on token overlap and may not align with human judgment. |
| Approach: | They propose an open-source model that assesses the question answering abilities of multimodal large language models. |
| Outcome: | Experiments show that the ACE-M3 model performs better than existing models and is more reliable than existing metrics. |
Copied to clipboard
| Challenge: | Existing models for question generation suffer from lack of diversity and bad sentence structures. |
| Approach: | They propose a framework that integrates flexible templates with a neural-based model to generate diverse expressions of questions with sentence structure guidance. |
| Outcome: | The proposed framework generates diverse expressions of questions with sentence structure guidance while maintaining high quality and consistency under automatic evaluation and human evaluation. |
Copied to clipboard
| Challenge: | generative models are less practical for building real-time conversation systems due to high latency and large memory footprint. |
| Approach: | They propose a method that preserves the efficiency of a retrieval model while leveraging the conversational ability of generative models. |
| Outcome: | The proposed method preserves the efficiency of a retrieval model while leveraging the conversational ability of generative models. |
Copied to clipboard
| Challenge: | Existing approaches to augment textual dialogues with retrieved images pose privacy, diversity, and quality constraints. |
| Approach: | They propose a framework to augment text-only dialogues with diverse and high-quality images by using a diffusion model and a feedback loop. |
| Outcome: | The proposed framework is comparable to or better than baselines, with significant improvements in human evaluation, especially against retrieval baselines where the image database is small. |
Copied to clipboard
| Challenge: | Answering non-factoid questions (NFQs) is a challenging task, requiring passage-level answers that are difficult to construct and evaluate. |
| Approach: | They propose a multi-document NFQA benchmark built on WikiHow, a website dedicated to answering “how-to” questions. |
| Outcome: | The proposed framework includes 11,746 human-written answers along with 74,527 supporting documents. |
Copied to clipboard
| Challenge: | Existing methods for evaluating RAArg are costly and lack long, complex arguments and real-world evidence. |
| Approach: | They propose to use multiple fine-grained LLM judges to evaluate RAArg using a new benchmark that features long and complex human-authored arguments on debated topics. |
| Outcome: | The proposed methods provide better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have given rise to generative recommenders . however, improving the generated content through user feedback is prohibitively expensive . |
| Approach: | They propose a generative explore-exploit method that exploits items with high engagement and actively explores hidden population preferences to improve recommendation quality. |
| Outcome: | The proposed approach exploits items with high engagement and actively explores hidden population preferences to improve recommendation quality. |
Copied to clipboard
| Challenge: | Persuasion dialogue systems have long-standing problems of dialogue repetition and inconsistency which could impact user experience and impede the persuaded outcome. |
| Approach: | They propose to refine a language model baseline without user simulators and distill sentence-level information about repetition, inconsistency, and task relevance through rewards. |
| Outcome: | The proposed model outperforms state-of-the-art models on automatic metrics and human evaluation results on a donation persuasion task and generates more diverse, consistent and persuasive conversations according to user feedback. |
Copied to clipboard
| Challenge: | a recent study shows that noisy reference summaries can be detrimental to model performance. |
| Approach: | They propose to selectively re-write unsupported reference sentences to better reflect source data. |
| Outcome: | The proposed method improves reference quality while retaining all data. |
Copied to clipboard
| Challenge: | Existing review summarization systems generate summary only based on review content and neglect the authors’ attributes (e.g., gender, age, and occupation). |
| Approach: | They propose an Attribute-aware Sequence Network (ASN) to take the aforementioned users’ characteristics into account by encoding their attributes over the words. |
| Outcome: | The proposed model outperforms existing systems on tripAtt and human evaluation by taking the authors' attributes into account and incorporating attribute embedding and word-using habits into word prediction. |
Copied to clipboard
| Challenge: | Existing studies on agreement-oriented multidocument summarization have focused on clusters of articles . a recent study focused on the use of a pretraining framework to summarize articles based on the "union" of the articles. |
| Approach: | They propose to use agreement-oriented multidocument summarization to provide agreement-orientated summaries that represent information common to all articles. |
| Outcome: | The proposed task is called agreement-oriented multidocument summarization . the authors apply the pretrained model PEGASUS onto the task . |
Copied to clipboard
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
Copied to clipboard
| Challenge: | Puns add the challenge of fusing commonsense and world knowledge with the ability to interpret lexical-semantic ambiguity. |
| Approach: | They propose to augment existing datasets with detailed crowdsourced annotations of puns, keywords and fine-grained funniness ratings to challenge current models' ability to understand and generate humor. |
| Outcome: | The proposed tasks include explanation generation to aid with pun classification and keyword-conditioned pun generation to challenge state-of-the-art models' ability to understand and generate humor. |
Copied to clipboard
| Challenge: | Existing evaluation models fail to identify lexical matching failures for open-domain question answering. |
| Approach: | They manually evaluate open-domain QA models by manually evaluating their answers on a popular benchmark. |
| Outcome: | The proposed model performs better on NQ-open than existing models and more than 50% of lexical matching failures are attributed to semantically equivalent answers. |
Copied to clipboard
| Challenge: | Synthetic translations have been used for a wide range of NLP tasks, but it remains unclear how they differ from naturally occurring data. |
| Approach: | They propose to use a semantic equivalence classifier to improve bitext quality without additional bilingual supervision to replace the originals. |
| Outcome: | The proposed samples improve bitext quality without additional bilingual supervision and are validated intrinsically and extrinsically through bilingual induction and MT tasks. |
Copied to clipboard
| Challenge: | Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient. |
| Approach: | They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics. |
| Outcome: | The proposed metrics show strong correlations with human judgments. |
Copied to clipboard
| Challenge: | Entity-centric summarization is a form of controllable summarizing that aims to generate a summary for a specific entity given a document. |
| Approach: | They propose to use a more abstract version of the original entity-centric ENTSUM summarization dataset to generate a shorter annotated summary for downstream users. |
| Outcome: | The proposed method is more abstract and uses supervised fine-tuning and large-scale instruction tuning to provide more specific and useful summaries for downstream users. |
Copied to clipboard
| Challenge: | Document grounded generation is the task of using the information provided in a document to improve text generation. |
| Approach: | They propose two new document grounded generation tasks that use information provided in a document to improve text generation. |
| Outcome: | The proposed models outperform existing methods on automated and human evaluation for closeness to reference and relevance to the document. |
Copied to clipboard
| Challenge: | Existing evaluation approaches to multi-document summarization of biomedical literature lack consistency and transparency. |
| Approach: | They propose a systematic approach to human evaluation of biomedical summaries and apply it to analyze the summary generated by two current evaluation models. |
| Outcome: | The proposed evaluation framework is based on two state-of-the-art models and examines the summaries generated by the two models to understand the deficiencies of existing evaluation approaches. |
Copied to clipboard
| Challenge: | Existing task-oriented dialog systems struggle to dynamically model long dialog context for interactions and effectively incorporate knowledge base (KB) information into dialog generation. |
| Approach: | They propose a dual dynamic memory network for multi-turn dialog generation . the model dynamically expands the dialog memory turn by turn and keeps track of dialog history . |
| Outcome: | The proposed model outperforms baseline models on three benchmark datasets on human evaluation and automatic evaluation. |
Copied to clipboard
| Challenge: | Consistency is a long standing issue faced by dialogue models. |
| Approach: | They propose to frame the consistency of dialogue agents as natural language inference and create a new natural language dataset called Dialogue NLI. |
| Outcome: | The proposed model can improve the consistency of a dialogue model with human evaluation and automatic metrics on a suite of evaluation sets designed to measure the model’s consistency. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks with limited references may not accurately reflect the quality of the model’s hypotheses. |
| Approach: | They propose a method to enrich evaluation benchmarks by diversifying the expression of a single reference into multiple high-quality ones to cover the semantic space of the reference sentence as much as possible. |
| Outcome: | The proposed method can enhance evaluation benchmarks by diversifying the expression of reference into multiple high-quality ones to cover the semantic space of the reference sentence as much as possible. |
Copied to clipboard
| Challenge: | Recent work has shown limited utility of natural language explanations in improving classification. |
| Approach: | They propose a two-stage few-shot learning framework that generates explanations and fine-tunes a smaller model with generated explanations. |
| Outcome: | The proposed framework increases inference accuracy over strong baselines, but human evaluation reveals that the majority of generated explanations does not adequately justify classification decisions. |
Copied to clipboard
| Challenge: | Low-resource languages such as those in the Finno-Ugric family are underrepresented in large language models. |
| Approach: | They propose to develop large language models for extremely low-resource languages . they focus on Vro, Livonian, and Komi, which are underrepresented . |
| Outcome: | The proposed models cover almost the entire cycle of creation, from data collection to instruction tuning and evaluation. |
Copied to clipboard
| Challenge: | Language modelling and machine translation tasks mostly use subword or character inputs, but syllables are rarely used. |
| Approach: | They explore the potential of syllables for open-vocabulary language modelling in 21 languages. |
| Outcome: | The proposed method outperforms characters and subwords in a non-related and low-resource language pair. |
Copied to clipboard
| Challenge: | a new method to detect political bias in news articles overcomes this domain dependency . partisan bias exists in various social issues, including the 2016 presidential election . |
| Approach: | They propose a multi-head hierarchical attention model that encodes the structure of long documents through a diverse ensemble of attention heads. |
| Outcome: | The proposed model outperforms existing methods for detecting political bias in news articles. |
Copied to clipboard
| Challenge: | a new dataset aims to automate the method to counter trolls . trolleds cause psychological damage to individuals and increase social costs . |
| Approach: | They propose to use a dataset to generate counter responses by varying counter responses according to a given strategy. |
| Outcome: | The proposed method improves strategy-controlled sentence generation. |
Copied to clipboard
| Challenge: | Existing work relies on commercial search engines and human evaluation, making it difficult to reproduce and compare different modeling approaches. |
| Approach: | They propose a new generation paradigm that requires large language models to provide citations to one or a few text passages for any statement they generate. |
| Outcome: | The proposed model improves factual correctness and verifiability of large language models by providing citations to a set of questions and retrieval corpora and generating answers with citation. |
Copied to clipboard
| Challenge: | Existing automated student answer assessment models lack explainable and faithful feedback. |
| Approach: | They propose a framework that leverages ChatGPT for student answer scoring and rationale generation. |
| Outcome: | The proposed method improves the overall QWK score by 11% compared to ChatGPT. |
Copied to clipboard
| Challenge: | Existing euphemism datasets are only domain-specific or language-specific. |
| Approach: | They propose a unified model to jointly conduct bilingual euphemism detection and identification tasks. |
| Outcome: | The proposed model is effective and provides a new reference standard for euphemism detection and identification. |
Copied to clipboard
| Challenge: | Existing datasets for supervised news summarization contain considerable amount of noise and expensive training data. |
| Approach: | They propose a large-scale and high-quality dataset for supervised abstractive news summarization containing 1.3 million training samples. |
| Outcome: | The proposed dataset is more factual and informative than established summarization datasets. |
Copied to clipboard
| Challenge: | Using the standard protocol to evaluate NLGs is often violated, resulting in annotator ratings cease to reflect their preferences. |
| Approach: | They propose a human evaluation protocol called system-level probabilistic assessment (SPA) this protocol is based on the assumption that annotators are biased by likert scales . |
| Outcome: | The proposed protocol can recover the ordering of GPT-3 models by size, but less than half of the expected preferences can be recovered when human evaluation is done with the standard protocol. |
Copied to clipboard
| Challenge: | Existing models for text-to-text generation do not explicitly focus on important concepts in the input and output. |
| Approach: | They propose a framework to automatically extract, denoise, and enforce important input concepts as lexical constraints. |
| Outcome: | The proposed framework performs comparably or better than its unconstrained counterpart on automatic metrics and receives better ratings in the human evaluation. |
Copied to clipboard
| Challenge: | Factual inconsistencies in generated summaries severely limit the practical applications of abstractive dialogue summarization. |
| Approach: | They propose a typology of factual errors to better understand hallucinations generated by current models and a contrastive fine-tuning strategy to improve the factual consistency and overall quality of summaries. |
| Outcome: | The proposed model significantly reduces all kinds of factual errors on both SAMSum dialogue summarization and AMI meeting summarizing datasets. |
Copied to clipboard
| Challenge: | Existing pre-trained summarization models produce text that is factually inconsistent with the input. |
| Approach: | They present a scale-based scale for Likert rating and a scoring algorithm for Best-Worst Scaling to improve crowdsourcing reliability. |
| Outcome: | The proposed model is more reliable than existing models on two news summarization datasets. |
Copied to clipboard
| Challenge: | Paraphrase generation is an important but challenging task in natural language processing . traditional symbolic approaches to paraphrase generation include rule-based methods, thesaurus-based approaches and statistical machine translation (SMT) |
| Approach: | They propose a deep reinforcement learning approach to automatic paraphrase generation . they propose supervised learning and reinforcement learning for evaluators . |
| Outcome: | The proposed framework outperforms state-of-the-art methods in paraphrase generation on two datasets. |
Copied to clipboard
| Challenge: | Long-form text generation remains a challenge for large language models . generating extended sequences often leads to degraded coherence and logical consistency . |
| Approach: | They propose a framework that integrates explicit structured thinking into long-form text generation. |
| Outcome: | The proposed framework surpasses even larger-scale models in evaluation and human evaluation. |
Copied to clipboard
| Challenge: | Multimodal summarization with multimodal output (MSMO) has attracted increasing research interest . evaluation is an emerging yet underexplored research topic . |
| Approach: | They propose a framework that studies three research questions of MSMO evaluation . they propose an automatic evaluation metric and a meta-evaluation benchmark dataset . |
| Outcome: | The proposed evaluation metric and human-annotated meta-evaluation benchmark are used to assess the quality of evaluation metrics and show the framework is effective. |
Copied to clipboard
| Challenge: | Unreliable evaluation guidelines can yield inaccurate assessment outcomes, potentially impeding the advancement of NLG in the right direction. |
| Approach: | They propose to collect annotated human evaluation guidelines and a method for detecting guideline vulnerabilities using Large Language Models. |
| Outcome: | The proposed dataset includes eight vulnerabilities and a method for detecting guideline vulnerabilities. |
Copied to clipboard
| Challenge: | Existing studies compare offline and online neural machine translation architectures . we examine the impact of online decoding constraints on the translation quality . |
| Approach: | They evaluate offline and online neural machine translation architectures using human evaluations on English-German and German-English language pairs. |
| Outcome: | The proposed models are particularly sensitive to latency constraints and are well-suited for offline translation tasks. |
Copied to clipboard
| Challenge: | Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed . |
| Approach: | They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues . |
| Outcome: | The proposed method is highly reliable while remaining feasible and low cost. |
Copied to clipboard
| Challenge: | Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy. |
| Approach: | They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method. |
| Outcome: | The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected. |
Copied to clipboard
| Challenge: | Existing studies have shown that non-autoregressive (NAT) methods underperform autoregressive methods (AT) however, their evaluation using BLEU has been shown to weakly correlate with human annotations. |
| Approach: | They propose to evaluate four representative NAT methods using BLEU to narrow the performance gap between autoregressive and autoregressive translations. |
| Outcome: | The proposed methods underperform NAT and autoregressive methods under more reliable evaluation metrics. |
Copied to clipboard
| Challenge: | Existing models for narrative story generation lack semantic dependency among sentences. |
| Approach: | They propose a skeleton-based model that generates the most critical phrases and expands them to a complete sentence. |
| Outcome: | The proposed model can generate significantly more coherent stories according to human evaluation and automatic evaluation. |
Copied to clipboard
| Challenge: | Existing generative search engines are rapidly gaining users, according to a new study . existing systems are poorly cited and lack reliability, a study finds . |
| Approach: | They conduct human evaluations of four popular generative search engines . they find that existing generative engines are fluent and appear informative . |
| Outcome: | The results show that existing generative search engines are not reliable and often contain unsupported statements and inaccurate citations. |
Copied to clipboard
| Challenge: | Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data. |
| Approach: | They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models. |
| Outcome: | The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective. |
Copied to clipboard
| Challenge: | CAT tools based on translation memories (TMs) are limited in their use for a number of translation tasks due to the limited availability of in-domain TMs. |
| Approach: | They propose a neural approach to exploit in-domain TMs and in-target-language (TL) monolingual corpora to exploit CAT tools. |
| Outcome: | The proposed approach exploits in-domain TMs and in-target-language (TL) monolingual corpora and increases translation proposals on four language pairs. |
Copied to clipboard
| Challenge: | Existing approaches to generalize multilingual dialogue systems to multilingual settings often make assumptions about data availability. |
| Approach: | They propose to transfer inductive biases for target languages learned by pretrained teacher models to student models via sequence-level knowledge distillation. |
| Outcome: | The proposed method performs well on the multiATIS++ benchmark, and is comparable to human annotations in both slot F1 and intent accuracy. |
Copied to clipboard
| Challenge: | Short-form video hashtag recommendation (SVHR) is a classification or ranking problem that selects hashtags from a set of limited candidates. |
| Approach: | They propose a short-form video hashtag recommendation task that better represents how hashtags are created naturally by retrieving relevant hashtags from a large-scale hashtag pool as extra guidance signals. |
| Outcome: | The proposed model outperforms strong classification baselines on two short-form video datasets and the guidance signals boost the performance by 8.11 and 2.17 absolute ROUGE-1 scores on average. |
Copied to clipboard
| Challenge: | Abstractive dialogue summarization aims to convert long dialogue content into its short form where the salient information is preserved while the redundant pieces are ignored. |
| Approach: | They propose to have the model perceive the redundant parts of an input dialogue history during the training phase. |
| Outcome: | The proposed method significantly outperforms baselines on the semantic matching and factual consistent based metrics. |
Copied to clipboard
| Challenge: | Existing models for dialogue summarization focus on document summarizing on time and speaker-centered points, but this approach is limited in understanding the dialogue. |
| Approach: | They propose a 2D view of dialogue based on a time-speaker perspective where the time and speaker streams of dialogue can be obtained as strengthened input. |
| Outcome: | The proposed model outperforms existing models on the QMSum dataset and improves summary faithfulness and human evaluation. |
Copied to clipboard
| Challenge: | Existing methods for evaluation of large language models are inefficient and inefficient due to inaccuracy of standard metrics in human perception of text quality and inefficiency in sampling informative test examples. |
| Approach: | They propose a sample-efficient human evaluation method for large language models based on the principle of MAximum Discrepancy (MAD) competition. |
| Outcome: | The proposed method achieves the “golden” ranking of LLMs with a minimum set of input instructions, which in turn reveal their relative strengths and weaknesses. |
Copied to clipboard
| Challenge: | Recent years have brought about interest in the task of summarizing conversation threads. |
| Approach: | They develop an email thread summarization dataset that contains human-annotated short and long email threads over a wide variety of topics. |
| Outcome: | The proposed dataset contains human-annotated short (30 words) and long (100 words) summaries of 2,549 email threads over a wide variety of topics. |
Copied to clipboard
| Challenge: | evaluators using large language models face ambiguous criteria and inconsistent evaluations. |
| Approach: | They investigate whether checklists should be used for all questions or selectively . they generate checklists using six methods and evaluate their effectiveness across eight models . |
| Outcome: | The proposed method improves evaluation performance in pairwise comparisons while ignoring human-written criteria. |
Copied to clipboard
| Challenge: | a new approach to contentful neural conversation is proposed . end-to-end models are effective in learning fluent responses, but their responses are often vacuous and uninformative. |
| Approach: | They propose a model that provides the conversation model with relevant text on the fly as a source of external knowledge. |
| Outcome: | The proposed model improves the informativeness and diversity of generated output compared to previous methods. |
Copied to clipboard
| Challenge: | Existing dialog inpainting methods generate ConvQA datasets with low contextual relevance due to insufficient learning of question-answer alignment. |
| Approach: | They propose a dialog inpainting method that generates ConvQA datasets from documents . they propose re-ranking tasks and a framework that generate contextually relevant questions . |
| Outcome: | The proposed framework generates ConvQA datasets with high contextual relevance from textual sources. |
Copied to clipboard
| Challenge: | Recent research has focused on literary machine translation (MT) but evaluation of literary MT remains an open problem. |
| Approach: | They propose a paragraph-level parallel corpus containing verified human translations and 13k evaluated sentences across four language pairs. |
| Outcome: | The proposed corpus compares human evaluations with students and professionals . it shows that the adequacy of human evaluation is controlled by two factors . |
Copied to clipboard
| Challenge: | Instruction-following LLMs have recently allowed systems to discover hidden concepts from a collection of unstructured documents based on a natural language description of the purpose of the discovery (i.e., goal). |
| Approach: | They propose a goal-oriented latent factor discovery system that integrates LLM’s instruction-following ability with statistical models to handle large, noisy datasets where LLM reasoning alone falls short. |
| Outcome: | The proposed system improves task performance by 5-52% over baselines and 1.8 times as often as the best alternative, on average, in human evaluation. |
Copied to clipboard
| Challenge: | In Natural Language Interfaces to Databases systems, text-to-SQL parsers allow users to query databases by using natural language questions. |
| Approach: | They propose a parser-independent interactive approach that interacts with users using multi-choice questions and can easily work with arbitrary parsers. |
| Outcome: | The proposed approach improves performance with limited interaction turns by using simulation and human evaluation on two cross-domain datasets with five state-of-the-art parsers. |
Copied to clipboard
| Challenge: | Using reinforcement learning to learn dialogue policy requires a large volume of interactions with users. |
| Approach: | They propose a task-oriented dialogue agent that efficiently learns dialogue policy from demonstrations . they use an imitation model to distill knowledge from demonstration and reward shaping . |
| Outcome: | The proposed agent efficiently learns dialogue policy from demonstrations through policy shaping and reward shaping. |
Copied to clipboard
| Challenge: | Moderation is essential for maintaining and improving the quality of online discussions. |
| Approach: | They annotate a dataset on 13 modes of discussion and use it to generate positive moderation. |
| Outcome: | The proposed model shows that professional moderation generates higher ratings than professional moderated moderation, but prefers professional moderate in pairwise comparison. |
Copied to clipboard
| Challenge: | Several studies use different information as ”pivot” such as language, semantic representation and so on. |
| Approach: | They propose to use visual information as the "pivot" of back-translation to generate paraphrases using paired image-caption data. |
| Outcome: | The proposed model generates paraphrase with good relevancy, fluency and diversity . it is based on paired image-caption data and can train a paraphrasing model . |
Copied to clipboard
| Challenge: | Sign words are the building blocks of any sign language. |
| Approach: | They propose a word-conditioned 3D American Sign Language (ASL) generation model that synthesizes real-time motion sequences for sign words. |
| Outcome: | The proposed model outperforms the baseline model in the task of sign word generation. |
Copied to clipboard
| Challenge: | Recent years have seen several shifts in summarization research, including extractive models. |
| Approach: | They propose a pipeline method for applying GPT-3.5 to summarize user reviews . they propose three new metrics targeting faithfulness, factuality, and genericity . |
| Outcome: | The proposed methods perform well in opinion summarization, the authors show . they also show that standard evaluation metrics do not reflect this performance . |
Copied to clipboard
| Challenge: | Using a dataset for sequential procedural (how-to) text generation from images, we show that 61% of the users found our proposed model is better than the baseline model in terms of overall recipes. |
| Approach: | They propose a dataset for sequential procedural (how-to) text generation from images in cooking domain. |
| Outcome: | The proposed model achieves a METEOR score of 0.31, an improvement of 0.6 over the baseline model. |
Copied to clipboard
| Challenge: | Currently, most reinforcement learning methods for dialog policy learning train a centralized agent that selects a predefined joint action concatenating domain name, intent type, and slot name. |
| Approach: | They propose a hierarchical multi-agent framework in which each part of the action is led by a different agent and a joint optimization process that makes agents can exchange their policy information. |
| Outcome: | The proposed framework reduces labor costs for action templates and decreases the size of the action space for each agent. |
Copied to clipboard
| Challenge: | Recent advances in end-to-end neural networks-based approaches have shown wide success in sequence generation tasks. |
| Approach: | They propose to optimize multiple metric rewards simultaneously using a multi-armed bandit approach . they empirically show the effectiveness of their approaches via various automatic metrics and human evaluation . |
| Outcome: | The proposed approach improves on question generation and data-to-text generation using a bandit approach. |
Copied to clipboard
| Challenge: | Text simplification is a valuable technique, but research on it is limited. |
| Approach: | They propose a document-level simplification task using Wikipedia dumps as a dataset and propose an automatic evaluation metric called D-SARI. |
| Outcome: | The proposed metric is more suitable for document-level simplification task. |
Copied to clipboard
| Challenge: | researchers have posited Dungeons and Dragons as a challenge problem to test systems on various language-related capabilities. |
| Approach: | They frame Dungeons and Dragons specifically as a dialogue system challenge . they train a large language model to generate the next game turn, conditioning it on different information. |
| Outcome: | The proposed game generates the next conversational turn and predicts the state of the game given the dialogue history. |
Copied to clipboard
| Challenge: | Automated metrics have reported flaws when applied to measure quality aspects of generated text and have been shown to correlate poorly with human judgements. |
| Approach: | They propose an agent-based framework to measure the required number of human annotations when evaluating generated outputs in relative comparison settings. |
| Outcome: | The proposed model can be compared with a crowdsourced case study and a simulation with simulated human judgements. |
Copied to clipboard
| Challenge: | Current factuality metrics do not account for vision modality, thus are not adequate for vision-and-language summarization. |
| Approach: | They propose a weighted combination of CLIPScore and BERTScore to evaluate factuality for abstractive document summarization. |
| Outcome: | The proposed metric outperforms existing factuality metrics on four factuity metric-evaluation benchmarks and is robust to human judgments. |
Copied to clipboard
| Challenge: | Multilingual NMT is an attractive solution for production, but to match bilingual quality, it comes at the cost of larger and slower models. |
| Approach: | They propose to use a shallow decoder with vocabulary filtering to speed up inference . they validate their findings with BLEU and chrF on 380 language pairs . |
| Outcome: | The proposed approach can be used in two 20-language multi-parallel settings. |
Copied to clipboard
| Challenge: | Scientific peer review is essential for the quality of academic publications. |
| Approach: | They propose a method that summarises scholarly reviews using a Rational Speech Act framework and novel uniqueness scores. |
| Outcome: | The proposed method generates more discriminative summaries than baseline methods in terms of human evaluation while achieving comparable performance with these methods in term of automatic metrics. |
Copied to clipboard
| Challenge: | Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs. |
| Approach: | They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria. |
| Outcome: | The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning. |
Copied to clipboard
| Challenge: | Existing studies focus on image-sharing behavior in singular sessions, leading to limited long-term social interaction. |
| Approach: | They propose a large-scale long-term multi-modal dialogue dataset that generates long-time multi-modity dialogue distilled from ChatGPT and proposed image aligner. |
| Outcome: | The proposed framework generates long-term multi-modal dialogue from ChatGPT and image aligner. |
Copied to clipboard
| Challenge: | Existing methods for estimating the effects of text on human evaluation are limited to testing a small number of pre-specified text treatments. |
| Approach: | They propose a method for flexibly discovering clusters of similar text phrases that are predictive of human reactions to texts using convolutional neural networks. |
| Outcome: | The proposed method can detect and predict human reactions to texts under certain assumptions. |
Copied to clipboard
| Challenge: | TuringQ is the first benchmark designed to evaluate the reasoning capabilities of large language models (LLMs) in the theory of computation. |
| Approach: | They propose a benchmark to evaluate the reasoning capabilities of large language models in the theory of computation. |
| Outcome: | The proposed system shows competitive accuracy when compared to human evaluation. |
Copied to clipboard
| Challenge: | a new benchmark summarization model is being developed to train few-shot summarizers . a large number of summarizing tasks are required to perform well in heterogeneous datasets. |
| Approach: | They propose a few-shot summarization model pre-trained with multiple summarizing tasks . they propose 'uniSumm' to be prefix-tuned to excel at any few-shot summarisation task . |
| Outcome: | The proposed model outperforms baseline models under automatic and human evaluations and achieves comparable results in human evaluation. |
Copied to clipboard
| Challenge: | a gap between conversations can be weeks, months or years, and dialogue systems which do not explicitly model time may generate unnatural responses. |
| Approach: | They propose to model the passage of time between conversations by exposing time information to a multi-session dialogue dataset and comparing different representations of time and event progress. |
| Outcome: | The proposed model is based on a real-time dataset showing that it can predict topics and information gained from conversations over a long time span. |
Copied to clipboard
| Challenge: | Experimental results show that the distilled language model outperforms its teacher model (ChatGPT) in most cases. |
| Approach: | They propose a Large Language Model (LLM) that leverages both distilled data from **ChatGPT** and real-world data from**doctors** in the supervised fine-tuning stage. |
| Outcome: | The proposed model outperforms the teacher model in most cases by using additional real-world data and RLMF to align the language model with the merits of both sources. |
Copied to clipboard
| Challenge: | Question Answering (QA) tasks require a mix of relevant and irrelevant information in these contexts to perform well. |
| Approach: | They propose a context filtering approach that removes non-essential details, summarizing crucial content through Reward Modeling. |
| Outcome: | The proposed approach outperforms baseline models in 6.8-folds. |
Copied to clipboard
| Challenge: | ATOMIC is a large-scale commonsense knowledge graph (CSKG) containing everyday if-then knowledge triplets, i.e., head event, relation, tail event. |
| Approach: | They propose a CSKG completion method called Rel-CSKGC to predict the relation given the head event and tail event of a triplet and train a model based on existing triplets. |
| Outcome: | The proposed method is based on existing triplets and can be used to complete the missing links in ATOMIC. |
Copied to clipboard
| Challenge: | telemedicine is a medical practice that provides patient care remotely using video conferencing tools. |
| Approach: | They build large-scale medical dialogue datasets to facilitate research . they pretrain several models on the Chinese MedDialog dataset and compare their performance . |
| Outcome: | The proposed datasets show that models trained on MedDialog can generate doctor-like medical dialogues. |
Copied to clipboard
| Challenge: | Paraphrasing of offensive content is a better alternative to content removal, but supervised methods often retain a large portion of the offensiveness of the original content. |
| Approach: | They propose to use In-Context Learning (ICL) to generate usable offensive paraphrases by using large language models. |
| Outcome: | The proposed framework is better than supervised methods on human evaluation and lower toxicity by 76%. |
Copied to clipboard
| Challenge: | Large “instruction-tuned” language models depend heavily on human-written instruction data . this limited quantity, diversity, and creativity hinders the generality of the tuned model . |
| Approach: | They propose a framework for improving instruction-following capabilities of pretrained language models by bootstrapping off their own generations. |
| Outcome: | The proposed framework outperforms existing public instruction datasets by 5% . it generates instructions, input, and output samples, then filters invalid or similar ones . |
Copied to clipboard
| Challenge: | Existing approaches to multi-task learning suffer from interference among datasets or fail to effectively reuse knowledge and skills learned from other datasets. |
| Approach: | They propose a sparsely activated modular network with a well-rounded set of operators and instantiate each operator with an independent module. |
| Outcome: | The proposed model outperforms state-of-the-art supervised approaches on 4 datasets with only 10% training data thanks to the modular architecture and multi-task learning. |
Copied to clipboard
| Challenge: | a new method to model news coverage of local government is needed . we show that newsworthiness predictions can be useful for journalists seeking to keep abreast of local governments. |
| Approach: | They propose a method that explicitly models when and why stories get press attention . they use an annotated corpus of news articles to build models that predict if a policy item will get covered . |
| Outcome: | The proposed model outperforms retrieval-based methods with limited annotated data and language use between corpora. |
Copied to clipboard
| Challenge: | Existing studies on response generation focus on relevance and fluency, rarely paying attention to the focus. |
| Approach: | They propose a Focus-aware response generation method that takes the focus into consideration and optimizes a multi-level encoder and focal decoder to generate multiple candidate responses. |
| Outcome: | The proposed method generates candidate responses that correspond to different focuses and performs better on two orthogonal inquiry conversation datasets. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is acknowledged as a challenging multi-modal task for Machine Learning (ML). |
| Approach: | They propose an interpretable approach for graph-based Visual Question Answering . their model is designed to intrinsically produce a subgraph during the question-answering process as its explanation . |
| Outcome: | The proposed model outperforms existing explainable methods on a graph-based VQA dataset. |
Copied to clipboard
| Challenge: | a new question answering corpus in french is designed to educational domain . we propose more complex questions and can justify the answers on validated material . |
| Approach: | They propose a question answering corpus in French designed to educational domain . they propose to propose more complex questions and justify answers on validated material . |
| Outcome: | The proposed question answering corpus is designed to be useful in educational domain . it proposes more complex questions and can justify answers on validated material . the proposed corpus could be used in the education domain, but it's not yet ready for use . |
Copied to clipboard
| Challenge: | Existing agents struggle due to bounded rationality in human data, low adaptability to counterpart behavior, and limited strategic reasoning. |
| Approach: | They propose a framework for turn-level offer optimization based on two core principles: opponent modeling and Tit-for-Tat reciprocity. |
| Outcome: | The proposed framework outperforms baselines across diverse partner agents and validates through human evaluation. |
Copied to clipboard
| Challenge: | Currently, alignment learning requires significant human demonstrations and feedback from proprietary LLMs such as ChatGPT. |
| Approach: | They propose a framework that uses synthetic feedback to align large language models to human values without extensive human annotations and proprietary LLMs. |
| Outcome: | The proposed model outperforms open-source models on human-annotated demonstrations in alignment benchmarks. |
Copied to clipboard
| Challenge: | Current vocabulary adaptation approaches append the target domainspecific vocabulary (V DOMAIN) at the end of the PLM vocabulary. |
| Approach: | They propose a vocabulary adaptation scheme that appends a target domain-specific vocabulary (V DOMAIN) at the end of the PLM vocabulary. |
| Outcome: | The proposed approach improves by 3.57% (in terms of accuracy) and 1.87% (royal-L) over various classification and summarization tasks. |
Copied to clipboard
| Challenge: | Human evaluation is indispensable for assessing the quality of texts generated by machine learning models or written by humans. |
| Approach: | They propose to use large language models to evaluate unseen texts using the same instructions and samples . they also use LLMs to generate responses to questions that are used to conduct human evaluation . |
| Outcome: | The proposed model can be used to evaluate texts in open-ended story generation and adversarial attacks. |
Copied to clipboard
| Challenge: | Existing studies have shown that social media users' posts can help identify depression, bipolar disorder or self-harm. |
| Approach: | They propose a hybrid abstractive summarisation approach combining hierarchical VAEs with LLMs to produce clinically meaningful summaries from social media timelines. |
| Outcome: | The proposed approach produces clinically meaningful summaries from social media user timelines, suitable for mental health monitoring. |
Copied to clipboard
| Challenge: | Prior work on document generation has tackled the creation of each separate format as a different task, leading to fragmented learning processes, redundancy in models and methods, and disjointed evaluation. |
| Approach: | They propose a method that unifies the generation and evaluation of templatic views of documents in multiple formats. |
| Outcome: | The proposed method improves performance for heterogeneous downstream applications while reducing the need for task specific evaluation metrics. |
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |
Copied to clipboard
| Challenge: | Quantization is widely used to improve inference speed and deployment of large language models. |
| Approach: | They conduct a thorough analysis of quantized multilingual LLMs . they find language disparately affected by quantization, non-Latin script languages worst . authors urge consideration of multilingual performance as evaluation criterion for efficient models . |
| Outcome: | The results show that quantization has harmful effects on human evaluation . language performance is disparately affected by quantization, the authors say . |
Copied to clipboard
| Challenge: | evaluating large language models' output is difficult due to the high cost of human evaluation. |
| Approach: | They propose a family of foundational large autorater models that train on over 100 quality assessment tasks. |
| Outcome: | The proposed model outperforms models on 8 of 12 autorater benchmarks on 53 quality assessment tasks. |
Copied to clipboard
| Challenge: | Social norms fundamentally shape interpersonal communication. |
| Approach: | They propose a human-in-the-loop pipeline to synthesize a bilingual dyadic dialogue dataset with turn-by-turn annotations of social norms for Chinese and American cultures. |
| Outcome: | The proposed dataset is high-quality through human evaluation and compares with existing models. |
Copied to clipboard
| Challenge: | Existing studies on concept design using text-to-image models have enabled rapid ideation of novel visual concepts. |
| Approach: | They propose a framework for generating novel, functionally coherent designs based on desired affordances by decomposing concepts into parts and affordance . they also develop a curriculum learning scheme that fine-tunes T2I models to progressively learn affordance composition while maintaining visual novelty. |
| Outcome: | The proposed framework outperforms state-of-the-art models for novelty and functional coherence in human evaluation. |
Copied to clipboard
| Challenge: | LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. |
| Approach: | They revisit meta-evaluations of LLM evaluators under a setting that more closely aligns with practice by examining evaluers’ ability to distinguish test system pairs that are closer in capability. |
| Outcome: | The proposed meta-evaluation setting is significantly different from the use of human evaluations. |
Copied to clipboard
| Challenge: | Existing approaches to creating inclusive vision-language models rely on human annotators, making it labor-intensive and creating cognitive burdens. |
| Approach: | They propose a semi-automated framework for constructing cultural VLM benchmarks . they use an annotated sample of Korean culture to generate questions . |
| Outcome: | The proposed framework is based on a Korean culture dataset and shows that open-source models lag behind proprietary ones in understanding Korean culture. |
Copied to clipboard
| Challenge: | Large language models (LLMs) use pretraining to predict the subsequent word, but less-resourced languages are being overlooked. |
| Approach: | They propose to expand the MLLM vocabularies to enhance expressiveness and use bilingual data for pretraining to align the high- and less-resourced languages. |
| Outcome: | The proposed model outperforms existing models in qualitative analyses compared to Korean monolingual models. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have aimed to refine their capacity to accurately follow human instructions and navigate intricate scenarios. |
| Approach: | They propose a method that uses a set of instructions to translate English into Japanese and then generates Japanese instruction data using GPT-4. |
| Outcome: | The proposed method outperforms Japanese-Alpaca models in the evaluation benchmarks without human references. |
Copied to clipboard
| Challenge: | Current mathematical benchmarks focus on evaluating MLLMs’ problem-solving ability, yet there is a crucial gap in addressing more complex scenarios such as error detection. |
| Approach: | They propose to evaluate multimodal error detection by evaluating two sub-tasks error step identification and error categorization. |
| Outcome: | The proposed task evaluates MLLMs' ability to handle multimodal questions compared to text-only models. |
Copied to clipboard
| Challenge: | Argument generation with diverse perspectives is essential for fostering balanced discourse and mitigating bias. |
| Approach: | They propose a Perspective-aware Preference Optimization with Entropy Maximization framework for diverse argument generation. |
| Outcome: | The proposed framework enhances perspective diversity through preference optimization based on the constructed preference dataset . |
Copied to clipboard
| Challenge: | a new study examines the effectiveness of large language models and non-LLMs in multimodal intent detection . large-scale multimodal data integrations include text, audio, and visual inputs . |
| Approach: | They propose a framework to debias multimodal intent detection datasets by using human evaluation. |
| Outcome: | The proposed framework debiases the datasets and shows that mistral-7B outperforms most competitive models by approximately 9% on MIntRec-1 and 4% on MIndRec2.0. |
Copied to clipboard
| Challenge: | Large Language Models generate human-like text, making them unreliable for biomedical relation extraction tasks. |
| Approach: | They propose to use Large Language Models as judges to evaluate biomedical relation extraction . they propose structured output formatting for LLM-generated responses that helps LLMs improve their performance by 15%. |
| Outcome: | The proposed method improves LLM-Judges' performance by 15% . it is cheaper and more efficient than human evaluation metrics, the authors say . |
Copied to clipboard
| Challenge: | Prior evaluation pipelines fail to evaluate factuality of long-form LLMs due to inefficiency and costly human assessment. |
| Approach: | They propose a fast and strong evaluation pipeline that can evaluate factuality of long-form LLMs . they propose 'faStFact' to reduce cost of web searching and inference calling . |
| Outcome: | The proposed evaluation pipeline achieves highest alignment with human evaluation and efficiency among existing baselines. |
Copied to clipboard
| Challenge: | Existing methods to extract causal relationships from medical case reports are insufficient for capturing causal relationships of an entire case. |
| Approach: | They propose a task that generates a causal tree with the primary disease as the root and extracts causal relationships from a medical case report. |
| Outcome: | The proposed method outperforms the baseline method by 20.2 points in the human evaluation and introduces evaluation metrics that reflect clinician preferences. |
Copied to clipboard
| Challenge: | Large language models produce content lacking pedagogical depth when asked to generate lessons . |
| Approach: | They propose a framework that allows teachers to select content according to pedagogical intent and sequence topics so foundations precede applications. |
| Outcome: | The framework achieves 67.8% win rate in human evaluation and 79.6% in LLM-based evaluation against eight baselines. |
Copied to clipboard
| Challenge: | Mainstream research in natural language processing has focused on high-resource and modern languages. |
| Approach: | They propose a task-anchored benchmark for Manchu–Classical Chinese translation . they use a parallel corpus of 16,627 sentence pairs to evaluate the model . |
| Outcome: | The proposed benchmarks show that linguistic differences influence performance and broader language coverage facilitate low-resource transfer. |
Copied to clipboard
| Challenge: | Existing methods for Jupyter Notebooks focus on generating cell-level descriptions from code snippets or table outputs independently. |
| Approach: | They propose a task to generate personalized cell-level descriptions using code, tables, and user-written guidelines in Jupyter Notebooks. |
| Outcome: | The proposed task combines code, tables, and user-written guidelines with personalized descriptions to evaluate the performance of existing models. |
Copied to clipboard
| Challenge: | KAHAN leverages LLMs as domain experts to drive the analysis. |
| Approach: | They propose a knowledge-augmented hierarchical framework that extracts insights from raw tabular data. |
| Outcome: | KAHAN outperforms existing frameworks on financial reporting benchmarks on narrative quality and factuality. |
Copied to clipboard
| Challenge: | Existing methods for topic-controllable summarization are limited by their recurrent architectures and require modifications to the model's architecture for controlling the topic. |
| Approach: | They propose a new topic-oriented evaluation measure to automatically evaluate the generated summaries based on the topic affinity between the generated summary and the desired topic. |
| Outcome: | The proposed method achieves better performance compared to more complicated embedding-based approaches while also being significantly faster. |
Copied to clipboard
| Challenge: | Claims are often nuanced and cannot be clearly labeled as “true” or “false” . however, a claim can be dissected into integral aspects and sub-aspects that are individually easier to validate . |
| Approach: | They propose a retrieval-augmented generation-based framework for deconstructing nuanced claims . claim can be dissected into integral aspects and sub-aspects, which are easier to validate . |
| Outcome: | The proposed framework can be easily deconstructed into integral aspects and sub-aspects, which are easier to validate. |
Copied to clipboard
| Challenge: | a new benchmark evaluates the truthfulness of large language models (LLMs) based on imitative falsehoods. |
| Approach: | They propose a professionally translated extension of the TruthfulQA benchmark . it evaluates truthfulness in Basque, Catalan, Galician, and Spanish . |
| Outcome: | The proposed extension of the TruthfulQA benchmark evaluates truthfulness in Basque, Catalan, Galician, and Spanish. |
Copied to clipboard
| Challenge: | Effective protection of private information is essential for knowledge dissemination in sensitive domains such as medical and legal. |
| Approach: | They perform a comprehensive study of privacy risks in LM-based summarization using closed- and four-weight models of different sizes and families. |
| Outcome: | The proposed models show that they leak personally identifiable information in their summaries, compared to human-generated summary summators, which show significantly higher privacy protection levels. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, resulting in inconsistent or unreliable generated text. |
| Approach: | They propose a logit-based ensemble method to measure LLM consistency and propose to use it to evaluate human ratings of LLM reliability. |
| Outcome: | The proposed method matches the best-performing existing metric in estimating human ratings of LLM consistency. |
Copied to clipboard
| Challenge: | Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. |
| Approach: | They propose a multimodal framework that retrieves supporting evidence from a paper and assigns each claim an overstatement score. |
| Outcome: | The proposed framework retrieves supporting evidence from ICLR and NeurIPS papers and assigns each claim an overstatement score. |
Copied to clipboard
| Challenge: | Existing data-to-text benchmarks that do not involve content selection feature short input-output pairs designed for sentence or paragraph-level generation with reference texts spanning only a few dozen tokens. |
| Approach: | They propose a system that generates multi-paragraph outputs in English and Irish . they compare a multi-agent configuration against a single-task variant . |
| Outcome: | The proposed framework generates multi-paragraph outputs in English and Irish . human evaluation and LLM-as-a-judge score better in both languages . |
Copied to clipboard
| Challenge: | Existing psychological counseling datasets suffer from monolithic client personas, insufficient therapeutic depth, and a lack of process controllability. |
| Approach: | They propose a framework that evolves static counseling corpora into high-fidelity dialogues . they use a Client Profiler that pairs life scenarios with psychological personality archetypes based on client personality and stage progression . |
| Outcome: | The proposed framework achieves 61-91% win rates against domain-specific baselines in pairwise evaluation and the highest average score in human evaluation, indicating potential for real-world counseling. |
Copied to clipboard
| Challenge: | Existing representations of hallucinations limit the types of errors that can be expressed, so we propose a new representation based on free-form textual descriptions, capturing the full range of possible errors. |
| Approach: | They propose a benchmark for localizing hallucinations using LLMs with a human annotation of over 1,000 examples and a protocol to verify its quality in a humans evaluation. |
| Outcome: | The proposed representation captures the full range of possible errors, and the best model achieves an F1 score of 0.67. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated strong performance in translation tasks. |
| Approach: | They propose a method that expands source-side data by rewriting original subtitles using information that can be extracted from the context, such as character profiles and scene descriptions. |
| Outcome: | The proposed method improves BLEU scores for film subtitle translation and achieves superior stylistic quality in human evaluation. |